The Journal of the Acoustical Society of America
● Acoustical Society of America (ASA)
Preprints posted in the last 90 days, ranked by how well they match The Journal of the Acoustical Society of America's content profile, based on 35 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Bozdogan, A.; Aarts, R. M.
Show abstract
Elephants and other large mammals produce low-frequency vocalizations extending well below the 20 Hz lower limit of human hearing, a regime known as infrasound. These rumbles serve vital social and reproductive functions over distances of several kilometers, yet they are inaudible to human observers and cannot be reproduced by conventional small loudspeakers. We present a complete signal-processing pipeline that renders sub-20 Hz elephant rumbles perceptible through a small loudspeaker by exploiting the missing-fundamental psychoacoustic effect. Butterworth bandpass filters isolate the infrasonic content; a full-wave integrator nonlinear device (NLD) generates the harmonic series required for virtual pitch perception; and a hysteresis-comparator fundamental-frequency estimator normalizes the NLD output. The pipeline was validated on African elephant field recordings and deployed on a credit-card-sized, low-cost single-board computer with an infrasound microphone and a small Bluetooth loudspeaker, demonstrating live operation in the field. The processed output shows a 10 dB to 15 dB elevation in the loudspeakers efficient band during call segments compared with background. The system enables zoo visitors and wildlife observers to perceive elephant rumbles in real time, opening new avenues for behavioral studies and public engagement with animal communication.
Garcia Ruiz, T.; Sanes, D. H.
Show abstract
Many perceptual skills improve with a few days of training. However, weeks or months of practice may be required to reach a level of expertise on complex tasks (Watson, 1980). Here, we explored how gerbils attain expertise on a difficult task: amplitude modulation (AM) rate discrimination at very shallow AM depths, similar to the depths used during vocal communication. Using an appetitive Go-Nogo procedure, we first trained 6 gerbils to perform an AM discrimination task (Nogo: 4 Hz; Go: 4.25-10 Hz) at a depth of 0 dB (re: 100% depth). Animals were then trained to perform AM discrimination at successively shallower depths, from -3 to -18 dB, requiring an average of 5-10 days of practice to reach a performance metric of d[≥]1 for each depth. Finally, we determined that AM discrimination thresholds were nearly identical between 0 to -12 dB, and only slightly elevated at -15 dB. Improvements in performance were accompanied by a large reduction in response time during procedural learning, and a gradual reduction of response time during perceptual learning, even as AM depth became shallower (i.e., more difficult). The shallowest depth at which gerbils displayed peak performance on the AM discrimination task is similar to their lowest AM depth detection thresholds. These results suggest performance on challenging auditory perceptual tasks require prolonged practice, and is accompanied by increased automaticity (i.e., lower response time) that stabilizes once expertise is achieved.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Benecke, J.; Whitmer, W. M.
Show abstract
In conventional hearing-aid personalisation, clinicians cannot hear what their patients hear, and patients cannot often reliably detect or describe what they hear. Self-adjustment avoids this issue but requires user controls that adjust hearing-aid signal processing parameters to be effective, efficient and easy. In this study, we explored (a) the roles of interface complexity and stimulus type in the self-adjustment of hearing-aid gain, and (b) how well individuals can adjust one sound to match another to assess the same interfaces and stimuli. Adult hearing-aid users with mild to moderate symmetrical sensorineural hearing loss repeatedly adjusted the gain (a) to their preference from individual prescription (n = 41) and (b) to match their previous preferences from a random starting point (n = 32) using three interfaces representing different bass/mid/treble configurations and three stimuli (music, speech and speech-in-noise). The large interindividual variability in self-adjusted gains clustered into three patterns of deviation from initial prescription: increased relative bass, overall gain reduction, and close to initial prescription. There were no substantial effects of interface nor stimulus on self-adjustment reliability (median {sigma} = 2.8 dB), whereas absolute sound-matching error increased with increasing interface complexity and centre frequency. Neither individual matching accuracy nor questionnaire responses predicted either self-adjusted gains or reliability. Overall, these results show that many - but not all - hearing-aid users can adjust gains with reasonable reliability, and while it can be difficult to predict the behaviour from the individual, the individual applies a similar self-adjustment behaviour across different interfaces and stimuli.
Teng, S.; Fusco, G.; Patel, A.
Show abstract
Blindness imposes constraints on the acquisition of environmental sensory information. To mitigate those constraints, some blind people employ active echolocation, a technique in which self-generated tongue "clicks" produce informative reflections from surrounding surfaces. Practitioners typically produce multiple clicks that guide, and are in turn shaped by, goal-relevant action. What perceptual information is gained in the echoacoustic signal from each click, and how does it inform motor behavior during task performance? To explore these poorly understood dynamics, here we recorded head movements and clicking behavior of an early-blind expert echolocation practitioner who localized and oriented toward a target object positioned at a 1 m distance and random azimuth in the frontal hemifield. Three additional participants, including a blind self-reported echolocator, were unable to perform the task better than chance level. Performance clearly benefited from available echoacoustic information: The larger target was localized with an average absolute angular error of 9.5{degrees} in 9.3 s, vs. 24.6{degrees} in 23.2 s for the smaller target. In a passive control condition prohibiting clicks altogether, no significant convergence on the target was observed, confirming the necessity of active sampling. Clicks were emitted somewhat more rapidly and intensely for small targets, but within-trial emission rate and head kinematics (left-to-right reversals) remained relatively invariant. Angular convergence toward the target was consistent with an exponential decay profile, though only weakly distinguishable from a linear trend for small targets. Pooled across trials within each condition, clicks were unimodally distributed about the target azimuth, suggesting an intensity-maximization strategy. In sum, clicking behavior and target size (therefore sonar strength) strongly influenced the rate and precision of orientation convergence toward the target, suggesting that dynamic interactions between motor-driven head movements, click production, and the resulting echoacoustic feedback accumulate goal-relevant evidence across multiple samples. Together, these results illustrate naturalistic sensorimotor dependencies underlying auditory active sensing in the absence of vision.
Hajicek, J.; Harris, S. E.; Neely, S. T.
Show abstract
PurposeThis research sought to develop a low-cognitive-load speech-in-noise test based on consonant confusions with the potential for assessing hearing-aid benefit. MethodsVowel-consonant-vowel (VCV) stimuli with added speech-shaped noise were presented as a closed-set consonant identification task. Initially, consonant-confusion matrices were used to select, from a larger set of consonants and vowel contexts, a set of ten consonants and associated signal-to-noise ratios (SNR) that were sensitive to hearing loss. The sensitivity of the qVCV test to hearing loss was validated by comparing predicted pure-tone average (PTA) hearing thresholds with their audiometric PTA. Clinical viability of the qVCV test was assessed by comparisons to the QuickSIN test. Hearing-aid benefit was assessed by comparing test scores in unaided and aided conditions. ResultsThe consonants most sensitive to hearing loss were /b d g t k v z s [esh] n/ in the vowel context /[a]/. A cross-validated prediction of PTA had a mean-absolute error of 5.7 dB. The repeatability of qVCV at 50 trials was equivalent to the QuickSIN average of two lists. Hearing-aid benefit was quantified as a decibel reduction in hearing loss. ConclusionsqVCV and QuickSIN performed similarly when test times are equated. The advantages of qVCV include lower cognitive demand, fewer learning effects, and automated scoring. PTA predicted by qVCV which greatly exceeds audiometric PTA may indicate either cognitive deficits or cochlear neural degeneration. The qVCV quantification of hearing-aid benefit may have clinical value.
Parra Pena, J. A.; Sorolla, C.; Quinteros Veas, N. F.; Ibarra, E. J.; Alzamendi, G. A.; Peterson, S. D.; Weerathunge, H. R.; Guenther, F. H.; Zanartu, M.
Show abstract
Accurate modeling of laryngeal motor control is key to understanding typical and disordered voice production. However, traditional biomechanical plant models based on ordinary differential equations (ODEs) often involve high computational costs and numerical instabilities, limiting their use in real-time closed-loop control frameworks. This study evaluates feature-driven machine learning (ML) regressors, specifically Random Forest (RF), Multilayer Perceptron Neural Networks (NN), and Polynomial Regression (PR), as surrogate forward models mapping laryngeal motor inputs to fundamental frequency and sound pressure level. Training data were generated with two biomechanical vocal fold models: the extended body-cover and the triangular body-cover. Results demonstrate that ML surrogates reduce execution times from seconds to milliseconds (e.g., 2 ms for PR), enabling stable real-time tracking via inverse Jacobian control. While RF provides the highest accuracy, NN and PR offer smoother control signals and smaller memory footprints. A practical performance threshold was identified near N = 1,000 training samples, below which accuracy degraded substantially when models were trained from scratch. These findings support ML surrogates as efficient and adaptable alternatives to direct numerical simulation, providing a foundation for future subject-specific modeling through transfer learning in data-limited clinical scenarios.
Rotaru, I.; Geirnaert, S.; Heintz, N.; Bertrand, A.; Francart, T.
Show abstract
Selective auditory attention decoding (AAD) enables tracking which of multiple concurrent speakers a listener attends to and is a key building block for neuro-steered hearing devices. While AAD integrated in a closed-loop system with real-time neurofeedback (NFB) is hypothesized to improve decoding through neural adaptation and error-correction behaviour, the short-term behavioral and algorithmic impact of such a bilateral human-machine interaction remains poorly understood. Here we evaluated the effects of NFB on AAD accuracy and user experience in a single-session AAD paradigm with online NFB involving nineteen participants. They performed a selective listening task with enforced attention switches across four conditions: open-loop (OL), closed-loop with auditory gain feedback (CLA), closed-loop with visual feedback (CLV), and a condition with pseudo-auditory gain control (psCLA) decoupled from the participants individual neural activity. AAD was performed online using both subject-specific and subject-independent linear decoders on 5 s sliding windows, followed by Hidden Markov Model post-processing. Online analysis showed comparable decoding performance across all conditions. However, offline posthoc analysis using subject-independent decoders revealed that AAD accuracy in the CLA condition was significantly lower than in the OL baseline. Subjectively, participants reported that CLA was significantly more distracting and required higher switching effort. Crucially, a causal analysis of the psCLA condition found no robust evidence that higher audio gains inherently improve decoding accuracy. Our results demonstrate that within a single-session paradigm with rapidly varying feedback cues, auditory neurofeedback may degrade AAD performance by increasing cognitive load and distraction. These findings suggest that suboptimal feedback can impede rather than facilitate learning. We conclude that more accurate and stable decoders and longitudinal, multi-session training protocols are likely essential prerequisites for achieving beneficial neurofeedback effects in closed-loop auditory attention systems.
Labib, S.; Liu, J.
Show abstract
Transcranial focused ultrasound is an emerging noninvasive neuromodulation technique offering high spatial precision and deep penetration. However, in deep brain neuromodulation in mice, the skull base attenuates the signal, distorting the focal region and creating off-target peaks. This study presents a machine-learning-driven simulation framework to optimize a bowl-shaped phased-array transducer design for hypothalamic targeting and compares its performance with that of time-reversal phase conjugation and a single-element baseline. A computed tomography-based mouse head model was used for full-wave acoustic simulations with a fixed bowl geometry (10 mm aperture, 6 mm radius of curvature). Designs were evaluated across various parameters, including operating frequency (0.2-1.5 MHz), active element count (16, 32, 64, 128), and element diameter (300-550 m). The evaluation employed four metrics: the presence of a -3 dB focal region within the hypothalamic area, axial focal length defined by the -3 dB full-width at half maximum, focal fragmentation measured by the -3 dB blob count, and targeting displacement. Random Forest surrogate models were trained in simulation outputs and paired with the Non-dominated Sorting Genetic Algorithm II to reduce computational costs during multi-objective optimization. The forward-excitation-optimized phased-array design (0.73 MHz, 128 elements, 381 m element diameter) achieved a focal region at the hypothalamic target with a full width at half maximum of 0.67 mm, a blob count of 1, and a targeting displacement of 0.38 mm when placed 1 mm below the nominal position. Time-reversal phase conjugation further improved confinement and targeting (full width at half maximum: 0.59 mm; displacement: 0.37 mm). Limitations include reliance on a single mouse anatomy, and incorporating additional CT-derived anatomies should enhance generalizability across strains, ages, and sexes. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=96 SRC="FIGDIR/small/727023v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@62de15org.highwire.dtl.DTLVardef@e26e57org.highwire.dtl.DTLVardef@1ba4893org.highwire.dtl.DTLVardef@f2c77a_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIA CT-based acoustic simulation and machine-learning framework was developed to optimize bowl-shaped phased-array transducers for mouse hypothalamic tFUS neuromodulation. C_LIO_LIRandom Forest surrogate models coupled with NSGA-II efficiently identified optimized array designs across frequency, element count, and element diameter. C_LIO_LIThe optimized phased-array design produced a compact hypothalamic focus with submillimeter targeting displacement, with further confinement achieved using time-reversal phase conjugation. C_LI
Zogby, D. S.; Eddington, V. M.; Craig, E. C.; Kloepper, L. N.
Show abstract
Common terns (Sterna hirundo) are regionally threatened migratory seabirds that form large breeding colonies during the North American summer months. They are highly vocal and serve as important bioindicators of aquatic ecosystems. Historically, acoustic studies on colonial seabirds have proven difficult due to the dense aggregations of individuals and high rate of call overlap. However, as passive acoustic monitoring (PAM) becomes increasingly common for studying seabird colonies, quantitative descriptions of species vocalizations are needed to accurately interpret behavioral information from colony soundscapes and support automated analysis of large acoustic datasets. This study aims to quantify the vocal repertoire of adult common terns. We deployed AudioMoths to collect acoustic data at a tern colony on Seavey Island, New Hampshire, USA from across the breeding season. Using RavenPro, unique call types were identified through visual and aural inspection of the acoustic data in the spectrogram. For each call, we then extracted measurements of peak frequency (Hz), bandwidth 90% (Hz), syllable duration 90% (s), and total bout duration (s) to quantify the characteristics of each call type. Statistical analyses for acoustic parameters by call type were performed using Kruskal-Wallis tests, followed by post-hoc Dunn tests. Our results demonstrate that each call type is significantly different from another by at least one parameter, with the exception of the kek and kip/tjuk calls. These findings present the first quantitative analysis of common tern vocalizations for North America. By defining temporal and spectral characteristics for multiple call types, this work helps translate colony soundscape into biologically meaningful information about tern behavior and colony dynamics. These descriptions also provide key parameters for developing automated tools to detect and classify vocalizations in dense, noisy colonies. Integrating quantified vocal characteristics with PAM offers a promising approach for monitoring colony activity and behavior while minimizing disturbance relative to traditional methods.
Herche, J. L.; King, C. D.; Groh, J. M.
Show abstract
Calibration of sound localization behavior in species with mobile eyes requires not only accurate visual input but also accurate oculomotor signals across the lifespan. The recent discovery of eye movement-related eardrum oscillations suggest that oculomotor signals may be incorporated into auditory processing at the level of the ear. One inference of this discovery is that individual variation in such signals might be correlated with individual variation in sound localization accuracy. Here, we tested this hypothesis in humans with normal hearing. We discovered that there is considerable variation in the accuracy of sound localization (here, saccades to sounds) even in normal individuals: median horizontal errors ranged from 2-6{degrees}, and median vertical errors could be as large as 36{degrees}. We separated the subject pool into groups with "good" performance (median vectorial error < 8{degrees}) vs "poor" performance (median vectorial error > 10{degrees}) and evaluated their respective EMREOs. The EMREOs differed across the two groups in both horizontal and vertical dimensions, in how saccade amplitude vs. initial eye position was encoded, and across time with respect to the saccade. These results are consistent with the interpretation that EMREOs are associated with underlying processes that ensure the accuracy of sound localization. HIGHLIGHTSO_LIThe accuracy of eye movements to look at sounds varied across individuals, with median errors spanning a greater than 10-fold range. This range is surprising given that the participants passed screening for normal hearing. C_LIO_LI"Good" vs "poor" sound localizers exhibited differences in their eye movement-related eardrum oscillations (EMREOs) C_LIO_LIEMREOs differed in both horizontal and vertical sensitivity, for both saccade amplitude and initial eye position, and the differences varied in timing with respect to saccade onset. C_LIO_LIWe interpret the results under the theory that poor sound localization may be a consequence of poor eye movement encoding, without which linking visual and auditory space is likely inaccurate. C_LI
Devolder, P.; Deloche, F.; Thienpont, M.; Keppler, H.; Verhulst, S.
Show abstract
The middle ear muscle reflex (MEMR) and medial olivocochlear reflex (MOCR) are increasingly studied for their role in suprathreshold auditory processing. However, recording these reflexes in humans is potentially complicated by age-related (sub)clinical hearing loss and co-activation. This study investigates (1) the influence of age-related (sub)clinical hearing loss, (2) methodological differences between conventional and wideband MEMR techniques, and (3) how MEMR activation contaminates MOCR recordings. Three test groups were included: young normal-hearing adults, middle-aged normal-hearing adults, and middle-aged adults with audiometric hearing loss. Cochlear status and neural encoding was assessed using distortion-product otoacoustic emissions (DPOAEs) and envelope following responses (EFRs). MEMR recordings were compared using conventional tonal stimuli and wideband stimuli. MOCR was recorded at elicitor levels of 60 and 75 dB to evaluate MEMR co-activation. MEMR was related to age, suggesting sensitivity to subclinical cochlear damage. Wideband stimuli were beneficial as elicitor (noise vs. tone), while changing the probe stimuli added no significant benefit (click vs. tone). MOCR strength did not correlate with age-related subclinical hearing, suggesting that MOCR measurements may reflect efferent function relatively independently of afferent sensorineural status in audiometric normal hearing subjects. However, reliable recordings were challenging in participants with audiometric hearing loss due to poor OAE baselines. MEMR co-activation was detectable in the click response and could alter MOCR-induced suppression. These findings suggest that, in cases of normal hearing thresholds, MEMR amplitude may be a marker of subclinical cochlear damage and MOCR measurements may more specifically reflect efferent function. Clinical measurements can be improved using broadband stimuli, accounting for outer-hair-cell damage, and defining criteria for reflex co-activation.
Dewey, J. B.
Show abstract
Mammalian hearing depends on the active amplification of sound-evoked waves as they travel along the basilar membrane within the cochlea. This amplification is mediated by the outer hair cells (OHCs), which generate force to enhance the vibrations of the surrounding structures. While OHCs at a given location only amplify basilar membrane motion for a narrow frequency range, recent measurements show that the amplification of motions deeper within the organ of Corti is much more broadband. However, the extent to which this broadband amplification influences the motions that are most relevant to inner hair cell stimulation - i.e., at the organs apical surface - remains uncertain. Here, optical coherence tomography was used to demonstrate that OHCs nonlinearly amplify the motions near the top of the organ of Corti, including at the reticular lamina and tectorial membrane, over a wide frequency range in the mouse cochlear apex. Responses at all frequencies were physiologically vulnerable and grew compressively with stimulus level. Low-frequency responses also exhibited non-monotonic features that were due to interference between amplified motion and the underlying traveling wave. The data suggest that broadband amplification of motions at the top of the organ of Corti likely explains certain phenomena observed in auditory nerve responses.
Marrone, J. P.; Ziliak, M. C.; Bartlett, E. L.
Show abstract
Auditory brainstem responses (ABRs) are a core part of objective functional evaluations of hearing sensitivity and subcortical auditory transmission. Manual assessments of ABR waveforms are still a primary means by which thresholds and peak amplitudes and latencies are measured, which is time-consuming and prone to user variability. Automated methods have offered promising alternatives for ABR classification, but they have sometimes been limited in accuracy or robustness. Here, we developed and tested a supervised convolutional neural network (CNN) based ABR peak classifier that works across sound levels and sound frequencies that can be run quickly on a personal computer using single or dual-channel ABR inputs. For ABR peaks I, III, IV, and V, the classifier achieved over 95% accuracy. High accuracy was maintained even after noise-exposure causing temporary or permanent threshold shifts, and over 90% of peaks were within 0.041 ms (1 sample) of the manually identified peak. Only a few hundred samples were needed to train the network, making it widely amenable to smaller data studies or where the number of subjects or sessions may be low.
Sztandera, J.; Poole, K. C.; Shiell, M. M.; Picinali, L.; Chait, M.
Show abstract
Detecting changes in acoustic environments is essential for situational awareness. It remains unclear whether the spatial location of such changes modulates automatic orienting and arousal mechanisms. We measured pupil dilation, pupil dilation rate, and microsaccade rate while listeners (n=25) heard complex, spatialized auditory scenes rendered over headphones using individualized HRTFs. Participants were naive to the critical manipulation: the appearance of a new source from one of five locations: front, left, right, back, or above. A subsequent localization task assessed perceptual spatial uncertainty. Behaviorally-irrelevant source appearances elicited a cascade of ocular responses. Microsaccadic inhibition emerged from [~]85ms after change onset, and was broadly comparable across locations, suggesting a location-invariant early orienting response to auditory change. Pupil dilation rate increased from [~]200ms, followed by a phasic pupil dilation response from [~]400ms, indicating engagement of arousal-related systems. Pupil responses were modulated by source location: changes from front/left/right elicited larger dilation than changes from above, with back responses showing a similar but weaker reduction. Behavioral localization revealed substantial confusion for front/back/above locations. However, this did not mirror the physiological data, as front sources elicited pupil responses comparable to lateral sources. These findings demonstrate that complex auditory scene changes recruit oculomotor and autonomic systems even outside the focus of task relevance. They further suggest a dissociation between early, location-invariant attentional capture indexed by microsaccadic inhibition and later, location-sensitive arousal indexed by pupil dilation. Spatial biases in auditory situational awareness therefore appear to emerge after initial change detection, shaping arousal and behavioral performance rather than the earliest orienting response.
Koshe, A.; Sobhani Tehrani, E.; Jalaleddini, K.; Motallebzadeh, H.
Show abstract
Quantifying the diagnostic dispersion of inferred parameter distributions is a challenge in uncertainty-aware modeling. Scalar summaries such as credible interval width are topology-blind; fundamentally different posterior morphologies can yield identical scores, obscuring whether a parameter is precisely estimated or constrained to a range. We propose a Composite Certainty Framework that addresses this metric degeneracy by aggregating five complementary uncertainty metrics including interquartile range, standard deviation, full width at half maximum, Shannon entropy, and mass width. These metrics are aggregated through non-parametric Borda rank voting into a single, unitless consensus certainty score. Applied to a simulation-based inference pipeline for a finite-element model of the human middle ear tuned to cadaveric acoustic measurements, the framework reveals parameter-specific identifiability profiles invisible to any individual metric. It produces two actionable clinical thresholds: (1) the maximum tolerable measurement noise for reliable parameter recovery, and (2) the minimum simulation budget for posterior convergence. We demonstrated that no single metric captures all aspects of posterior dispersion, as spread-based metrics and entropy diverge systematically for clinically critical parameters, whereas their aggregation produces a consensus reflecting genuine diagnostic certainty. The framework is generalizable to any model-based diagnostic pipeline where posterior distribution not merely its coverage, but determines clinical certainty.
Guo, Z.-c.; McFarlane, K.; McHaney, J. R.; Choksi, I.; Feeney, M.; Preston, L.; Chandrasekaran, B.
Show abstract
ObjectivesObjective and ecologically valid measures of speech processing can complement conventional audiologic assessments. Phoneme-related potentials (PRPs), derived by averaging listeners electroencephalography (EEG) responses time-locked to phonemes in continuous speech, have emerged as a promising approach for capturing cortical processing of speech in naturalistic listening conditions. Importantly, PRPs reveal speech perception challenges even when conventional audiograms are clinically normal, positioning them as a promising neural marker for suprathreshold listening difficulties that standard audiometry often misses. As a critical step toward clinical translation, this study examined the extent to which PRP-derived measures remain stable across real-world contexts relevant to clinical implementation, including monaural versus binaural presentation, stimulus intensity level, and repeated testing sessions. The study also assessed cortical tracking of lower-level speech acoustics to determine whether the PRP findings could be attributed to acoustic processing. DesignEEG was recorded from 18 young adults with normal hearing as they listened to audiobook speech presented monaurally or binaurally at 60 or 75 dB across two sessions separated by approximately one week. Neural differentiation of phoneme manner-of-articulation classes (vowels, nasals/approximants, fricatives, and stops) in PRPs was quantified using two measures: an F-statistic reflecting between-manner relative to within-manner variability, and classification accuracy from a machine-learning model trained to predict manner class from PRPs. Temporal response function modeling assessed neural tracking of continuous acoustic envelope and onset features of the audiobook speech. ResultsNeither PRP-derived measure of manner differentiation showed significant effects of session, presentation modality, intensity level, or their interactions. Intraclass correlation analyses further indicated moderate-to-good reliability across all three factors. In contrast, neural tracking of the acoustic envelope and acoustic onsets was stronger under binaural than monaural presentation, with binaural presentation eliciting more pronounced cortical responses to the envelope. ConclusionsPRP-derived measures remained relatively stable across modest procedural variations that are common in clinical testing contexts, positioning PRPs as a potent objective index of naturalistic speech processing. This stability may reflect cortical processing of abstract, linguistically relevant speech categories and suggest that PRPs provide complementary information beyond audiologic assessments of peripheral auditory functions and EEG measures that primarily capture lower-level acoustic processing.
Bilger, H.; J. Ryan, M.; Clarke, J.
Show abstract
The human larynx, compared to those of closely related primates, lies deeper in the throat and lacks vocal membranes and air sacs. These shifts are usually analyzed regarding their acoustic effects on vowel-like vocalizations, since the evolution of speech was long thought to require an expansion of vocal range driven by vocal tract modifications. However, vowels are just one type of phoneme, and speech is just one class of human utterance. To understand the evolutionary underpinnings of known shifts in human vocal morphology, a broader bioacoustic comparison is needed. Specifically, the range of sounds used in human speech must be compared to that employed in other human vocalizations and in the repertoires of extant close primate relatives. Here, we measure the acoustic-feature space occupied by human speech, non-linguistic, and musical vocalizations along with the calls of chimpanzees, bonobos, and chacma baboons. We use Mel-frequency cepstral coefficients to create an acoustic space depicting the spectro-temporal features of over 750,000 brief vocal segments sourced from published databases and other verified sources. Speech and song occupied significantly less volume in this acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volumes of speech and song were not statistically distinct from those of non-human primates. These results suggest that speech was not enabled by an expansion of human vocal acoustic space. Anatomical shifts unique to humans may have led to an elaboration of non-linguistic utterances, but learned vocalizations use a surprisingly small fraction of this space. Our understanding of human vocal evolution will be further informed by additional systematic comparisons of the function and homology of non-speech vocalizations, along with the collection and incorporation of more complete non-human primate vocal datasets, especially from Gorilla and Orangutan.
Lien, J. T.-H.; Strahl, S.; Garcia, C.; Vickers, D.
Show abstract
The human auditory system decomposes complex sounds into distinct components via a collection of processing steps. Knowing whether Spiral Ganglion Cells (SGCs) play an active role in the decoding of complex sounds can facilitate the development of Cochlear Implant (Cl) coding strategies and clinical assessment tools. Early animal studies reported SGCs being similar across different characteristic frequencies (CFs). In this study, human electrically evoked compound action potentials (eCAPs) were analysed to probe the relationship between the reciprocal of CF and the duration of the eCAP. A significant relationship could indicate that SGCs may not simply be passive cables. eCAP datasets from 6 published studies (175 Cl users, 1243 recordings) were analysed and their peaks were automatically labelled. The nlp2 latency was derived for each recording as a proxy of the action potential duration. The CF of each recording was estimated by mapping the average insertion angle of the electrode to the human SGC map. A weak but statistically significant relationship was observed between the n1p2 latency and the reciprocal of CF (random-effects model with random intercepts for subject, r = 0.09, p = 0.024, n= 450) supporting the hypothesis that lower CF is associated with slower repolarisation (longer n1p2 latency) in human spiral ganglion cells.
Shannon, A. J.; Barton, D. A. W.; Homer, M.; Houghton, C. J.
Show abstract
Segregation of speech into syllables is a key step in neural speech processing. It relies on the alignment of neural activity with the rhythmic structure of speech. Two competing hypotheses explain this neural speech tracking, phase-resetting and evoked responses. While phenomenological modelling of these hypotheses has been successful, we still lack understanding of the underlying cortical circuits. To investigate these mechanisms, we evaluate whether a biophysical next-generation neural mass model can reproduce several features of neural speech tracking, using phenomenological models of the competing hypotheses as algorithmic baselines. We investigate the models dynamics with four tests: recreating in-silico an EEG experiment that identified a correlation between tracking strength and phoneme sharpness, computing the Phase Concentration Metric, testing the effect of varying syllabic rates, and evaluating the Inter Event Phase Coherence across phoneme onsets. While all of the models that we study reproduce the sharpness-tuned rhythmic speech tracking, the evoked model requires a pre-processed acoustic edge impulse stimulus. We demonstrate that the neural mass model is performing thresholded phase-resetting triggered by sharp onsets in the continuous speech envelope. This produces cross-frequency nested oscillations that qualitatively match an experimentally-observed dual-peak signature in the Inter Event Phase Coherence. Our results indicate that the biophysical neural mass model provides a mechanistic bridge between generic oscillatory dynamics in cortical populations and the cognitive computations of speech tracking. Indeed, the non-linear dynamics of the neural mass model offer an explanation for how peak-rate event representations in auditory cortex activity arise in response to continuous acoustic input. Significance StatementSyllable segregation is crucial but challenging as natural speech lacks clear boundaries, yet humans perform this computation effortlessly. Speech aligns neural activity to syllabic rhythms, predicting syllable timing, but the underlying cortical mechanisms remain unknown. Relating this macroscopic behaviour to neurobiology is challenging; however, next-generation neural mass models promise to resolve this. We demonstrate that these models reproduce sharpness-tuned tracking and acoustic edge extraction. Dynamical analyses indicate this occurs through thresholded phase-resetting to phoneme onsets, triggering cross-frequency nested oscillations. Our results both advance biophysical understanding of syllable segregation and validate the models capacity for simulating macroscopic neural activity. These models offer a bridge between the neurobiology of the auditory cortex and speech processing dynamics that phenomenological models cannot provide.